Papers by Daan van Esch

6 papers
Language ID in the Wild: Unexpected Challenges on the Path to a Thousand-Language Web Text Corpus (2020.coling-main)

Copied to clipboard

Challenge: Large text corpora are increasingly important for a wide variety of NLP tasks.
Approach: They propose to train automatic language identification models on up to 1,629 languages . they find that human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages.
Outcome: The proposed models achieve over 90% average F1 on 1,629 languages . human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages - suggesting a need for more robust evaluation.
Connecting Language Technologies with Rich, Diverse Data Sources Covering Thousands of Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing data sources for many thousands of languages are rich and diverse . Efforts are ongoing to extend technology to many more of the world's languages .
Approach: They provide an overview of some of the major online data sources available for thousands of languages.
Outcome: The proposed language technologies are based on the data available for thousands of languages.
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)

Copied to clipboard

Challenge: LinguaMeta is a unified repository of language metadata for thousands of languages.
Approach: They introduce LinguaMeta, a unified resource for language metadata for thousands of languages.
Outcome: The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages.
Text Normalization Infrastructure that Scales to Hundreds of Language Varieties (L18-1)

Copied to clipboard

Challenge: a multi-language text normalization infrastructure is used to train language models for keyboards and speech recognition systems.
Approach: They describe a multi-language text normalization infrastructure that prepares textual data to train language models used in Google's keyboards and speech recognition systems.
Outcome: The proposed system can normalize training data across hundreds of languages . it can detect errors in training data and detect corruption issues .
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)

Copied to clipboard

Challenge: Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties.
Approach: They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each.
Outcome: The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations